Go top
Paper information

Evaluación del razonamiento jurídico de los modelos de lenguaje ante el examen de acceso a la Abogacía Española

Evaluation of the Legal Reasoning of Language Models on the Spanish Bar Admission Exam

L.F. S. Merchante, O. Tejerina Rodríguez, A. Rodríguez Jaramillo, B. Aguayo Martínez‐Sagrera, M. de Prada Rodríguez, K.J. Castro Matute, J. Plaza García, M.E. de Paz Carmona

Derecho Digital e Innovación Nº. 25

Original summary:

La aplicación de modelos lingüísticos de inteligencia artificial en el ámbito legal plantea desafíos críticos. Aunque en inglés han mostrado gran rendimiento en lógica y argumentación jurídica, faltan evaluaciones en español. Surgen así dos preguntas: si esos resultados son extrapolables al castellano y si alcanzarían un nivel adecuado de razonamiento jurídico. Para explorarlo, se organizó un hackathon con estudiantes de Ingeniería y Derecho, quienes probaron a los LLMs frente al examen oficial de acceso a la abogacía en España. Los sistemas lograron alta precisión sin justificar respuestas (85,5/100), pero descendieron cuando se exigió razonamiento legal (66,67/100), variando según la materia. Se concluye que los LLMs funcionan mejor en inglés, requieren evaluaciones multidimensionales y aún precisan supervisión jurídica y ética para su uso profesional.


English summary:

The application of artificial intelligence linguistic models in the legal field poses critical challenges. Although they have performed well in English in terms of logic and legal argumentation, there is a lack of evaluations in Spanish. This raises two questions: whether these results can be extrapolated to Spanish and whether they would achieve an adequate level of legal reasoning. To explore this, a hackathon was organised with engineering and law students, who tested the LLMs against the official bar exam in Spain. The systems achieved high accuracy without justifying answers (85,5/100), but their performance declined when legal reasoning was required (66,67/100), varying according to the subject matter. It was concluded that LLMs work better in English, require multidimensional evaluations, and still need legal and ethical supervision for professional use.


Spanish layman's summary:

Se evaluaron modelos de lenguaje de IA con el examen de acceso a la abogacía en España. La mejor media fue 85,5/100 en aciertos y 66,67/100 al exigir una justificación jurídica correcta, lo que apoya una evaluación multidimensional y supervisión experta.


English layman's summary:

AI language models were tested on Spain’s bar admission exam. The highest average was 85.5/100 for answer accuracy and 66.67/100 when correct legal justification was required, supporting multidimensional evaluation and expert oversight.


Keywords: Modelos de lenguaje, LLM, abogacía, examen de acceso, inteligencia artificial; Language models, LLM, advocacy, entrance examination, artificial intelligence


Published on paper: September 2025.



Citation:
L.F. S. Merchante, O. Tejerina Rodríguez, A. Rodríguez Jaramillo, B. Aguayo Martínez‐Sagrera, M. de Prada Rodríguez, K.J. Castro Matute, J. Plaza García, M.E. de Paz Carmona, "Evaluación del razonamiento jurídico de los modelos de lenguaje ante el examen de acceso a la Abogacía Española", Derecho Digital e Innovación, Nº. 25, September 2025.

    Research topics:
  • Natural Language Processing and Generative AI
  • Safe, Trustworthy, Fair and Interpretable AI
  • Ethical considerations in technology and Artificial Intelligence
    Research groups:
  • Instituto de Investigación Tecnológica (IIT)
    ODS:
  • Goal 16: Peace, justice and strong institutions
  • Goal 9: Industry, innovation and infrastructure
  • Goal 4: Quality education

pdf Preview
Request Request the document to be emailed to you.